Papers with overcoming text-dominated unimodal reliance
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions (2026.acl-long)
Copied to clipboard
| Challenge: | Recent Spoken Language Models lack the capability to discern Third-Party Interruptions (TPI) from the primary user’s ongoing flow, leaving them vulnerable to contextual failures. |
| Approach: | They propose a dataset with speaker-aware hard negatives to enforce acoustic cue prioritization for interruption handling and a framework to measure the interruption-handling strategy and precise speaker discrimination in deceptive contexts. |
| Outcome: | The proposed framework mitigates semantic shortcut learning while neglecting acoustic signals essential for discerning speaker changes. |